Communications Medicine
○ Springer Science and Business Media LLC
Preprints posted in the last 90 days, ranked by how well they match Communications Medicine's content profile, based on 113 papers previously published here. The average preprint has a 0.14% match score for this journal, so anything above that is already an above-average fit.
Zubair, A.; Whitcroft, K.; Khong, G.; Bhargava, E.
Show abstract
Background: Olfactory dysfunction is a recognised but poorly characterised comorbidity of Primary Ciliary Dyskinesia (PCD). No prior systematic review has synthesised its prevalence or clinical correlates. Methodology: A PRISMA compliant systematic review and meta-analysis was conducted. Five databases were searched to February 2026. Observational studies reporting olfactory function in confirmed PCD were included. Risk of Bias was assessed using the Newcastle-Ottawa Scale. A random-effects meta-analysis using the Freeman-Tukey double arcsine transformation was performed to calculate pooled prevalence with 95% confidence intervals (CI) and prediction intervals (PI). Results: Twelve studies (n=865) were included. Overall pooled prevalence of olfactory dysfunction was 43.4% (95% CI 25.2-62.5%; 95% PI 0.1-99.0%). Objective psychophysical testing yielded a significantly higher pooled prevalence of 66.1% (95% CI 55.5-76.0%; 95% PI 38.4-88.9%) compared to patient-reported outcome measures (30.5%; 95% CI 11.4-54.0%). Older age, greater sinonasal disease burden, and specific ciliary ultrastructural defects were associated with worse olfactory function. A striking discordance between objective dysfunction and subjective awareness was observed across multiple studies. Conclusions: Olfactory dysfunction is highly prevalent in PCD and substantially under-recognised by patients. Routine objective olfactory screening should be integrated into standard multidisciplinary PCD care.
Martini-Stoica, H.; Rupp, B. T.; Kunz, M.; Livraghi-Butrico, A.; Okuda, K.; O'Neal, W.; Randell, S.; Dang, H.; Murano, H.; Furusho, M.; Morton, L.; Askin, F.; Thorp, B. D.; Klatt-Cromwell, C.; Ebert, C. S.; Senior, B. A.; Vuncannon, J. R.; Kimple, A. J.; Byrd, K. M.
Show abstract
Background: Juvenile nasopharyngeal angiofibroma (JNA) is a rare locally aggressive vascular sinonasal tumor that primarily affects adolescent males. Despite advances in endoscopic surgery and preoperative embolization, JNA can be associated with major operative bleeding risk and clinically meaningful recurrence, while non-surgical treatment options remain limited. Methods: To define the cellular programs underlying JNA vascularity, we performed single-cell RNA sequencing of JNA tumors (n=2), tumor-adjacent mucosa, and control sinonasal tissue. We analyzed cell composition, differential gene expression, pathway enrichment, and cell-cell communication, followed by Drug2cell-based mapping of transcriptional states to candidate therapeutic targets. Results: JNA contained an expanded fibrovascular compartment composed of endothelial cells, fibroblasts, pericytes, vascular smooth muscle cells, and neural crest-like cells. Neural crest-like cells were enriched in JNA but showed relatively limited transcriptional differences from tumor-adjacent tissue. By contrast, endothelial cells demonstrated the strongest disease-associated remodeling, with enrichment of angiogenesis, extracellular matrix organization, hypoxia response, and cell migration pathways. Endothelial cells also showed downregulation of adaptive immune signaling pathways, suggesting reduced immune engagement within the tumor microenvironment. Intercellular communication analyses revealed dense endothelial-stromal signaling across the JNA fibrovascular network. Drug2cell analysis nominated VEGF/VEGFR signaling as a candidate therapeutic vulnerability, with VEGFR-targeting agents predicted to act primarily on vascular and lymphatic endothelial populations. Conclusions: JNA is organized around an angiogenesis-dominant fibrovascular program driven by endothelial-centered signaling. These data support further investigation of VEGF/VEGFR-directed therapy as a potential adjunctive strategy for patients with recurrent, unresectable, or surgically high-risk JNA.
Ravoni, A.; Liu, Y.; Cairo, S.; Castiglione, F.; Nardini, C.
Show abstract
Hepatoblastoma (HB) is the most common pediatric liver cancer and represents a major clinical challenge, due to the lack of effective therapies for advanced stages and disease relapse. In this work, we use the results of a previously HB-tailored agent-based model of the immune system to investigate whether model-derived variables can be of use in the prediction of patients' outcomes. To this aim, we apply factor analysis to the results of a simulated cohort of HB patients, to identify combinations of key immunological variables able to discriminate disease outcomes in the simulator, and we then assess the coherence of such predictions with independent results of differential expression and enrichment analyses on HB transcriptomics. Our analysis proposes that the ability of immune cells, particularly natural killer and CD8+ cytotoxic T cells, to recognize tumor-associated antigens and exert cytotoxic activity is essential for disease control following treatment.
Parker, T. M.; Oermann, E. K.; Grossman, S. N.; Kenney, R. C.
Show abstract
Background: Artificial intelligence (AI) systems for glaucoma diagnosis and prognostication from visual fields (VF) are under active development, yet do not audit for vertical-meridian-respecting field loss - known sequelae of stroke, hemorrhage, and neoplasm. We developed a self-supervised encoder of automated perimetry that learns anatomically interpretable VF structure without labels, and evaluated its capacity to identify suspected neurologic VF patterns in an independent public glaucoma dataset. Methods: We pretrained a 128-dimensional masked autoencoder on 23,223 unlabeled Humphrey VFs (patient-grouped training split of 28,943 fields from 3,871 patients; UWHVF, all-comers perimetry), using monocular pattern-deviation input. A supervised linear classifier over vertical-midline latent dimensions was trained on per-eye expert neurological/non-neurological labels and assessed under hard-negative cross-validation, with specificity evaluated on 100 held-out, structurally separated UWHVF controls. External evaluation used the Harvard-Glaucoma Fairness dataset (Harvard-GF; 3,300 patients with paired VF and optical coherence tomography [OCT] from a single academic center), which contributed no data at any training stage. Results: Masked reconstruction recovered structure concordant with retinal neuroanatomy: 50 of 128 latent dimensions emerged spatially specialized, versus 23 for the total-deviation encoder. The classifier achieved cross-validated balanced accuracy 0.78 (95% CI, 0.75-0.82) and AUC 0.85 (95% CI, 0.82-0.89), with no false positives among the 100 held-out controls. Applied to Harvard-GF without fine-tuning, it identified a top-20 of 1,748 glaucoma-labeled patients (1.1%) with morphology inconsistent with glaucoma; all 20 were positive on the rule-based Neurological Hemifield Test (mean score 62.4), and OCT showed preserved superior (Cohen d = +0.68; P < .001) and inferior (d = +0.63; P = .003) retinal nerve fiber layer versus severity-matched controls. Conclusions: A self-supervised VF encoder learned anatomically interpretable visual field structure from unlabeled data and identified suspected neurological cases in a curated glaucoma dataset, with expert, rule-based, and OCT corroboration. Visual field datasets used to train glaucoma AI may benefit from neurological screening before model training; the encoder reported here supports such audits and provides a foundation for neuro-ophthalmic AI beyond fundus photography and OCT.
Kwon, S.; Lee, C. S.; Lee, A. Y.; Zhang, L.
Show abstract
Purpose: To evaluate whether fluorescence lifetime imaging ophthalmoscopy (FLIO) combined with deep learning can detect metabolic signatures for classification of type 2 diabetes mellitus (T2DM). Design: Cross-sectional analysis of participants included AI-READI dataset (version 3) with FLIO imaging and and hemoglobin A1c (HbA1c) measurement. Subjects: 1,783 participants from the AI-READI dataset (version 3) with HbA1c measurements and FLIO imaging scans (6,912 total): 671 normoglycemic, 726 prediabetic, and 386 diabetic. Methods: Mean fluorescence lifetime maps were generated using a center-of-mass approach and used as inputs to AI models. We trained convolutional neural networks (CNNs), ResNet-18, and XGBoost under three-class (normal, prediabetic, diabetic) and two binary (normal vs. impaired; normal vs. diabetic) classification schemes, using nested 5-fold cross-validation with participant-level grouping. Main Outcome Measures: Macro-averaged accuracy, F1 score, area under the receiver operating characteristic curve (AUROC), sensitivity, specificity, and positive predictive value (PPV). Results: Group-averaged lifetime maps demonstrated consistent spatial differences across glycemic groups, with progressively longer lifetimes from normal to diabetic participants. The CNN achieved the best overall performance in the 3-class classification (accuracy 0.41 +/- 0.03, F1 score 0.39 +/- 0.02, AUROC 0.58 +/- 0.02), compared to the random classifier for 3-class classification (AUROC = 0.50; accuracy = F1 = 0.33). ResNet-18 and XGBoost showed similar performance (AUROC 0.53-0.58). Confusion matrices revealed substantial overlap between classes, with frequent misclassification toward the prediabetes group. Binary reformulation (normal vs. diabetic) improved performance substantially, with the CNN resulting in AUROC 0.63 +/- 0.02 and XGBoost 0.67 +/- 0.07. Conclusions: FLIO-derived lifetime maps capture metabolic signals associated with glycemic status but yield modest classification performance with current AI models. These findings highlight both the potential and the challenges of using FLIO for early metabolic screening and monitoring, informing future development of clinically applicable imaging biomarkers.
Goroshchuk, O.; Koller, D.
Show abstract
Background: Endometriosis affects approximately 10% of reproductive-age women and is associated with substantial diagnostic delay and heterogeneous symptom presentation. Prior machine-learning prediction models have relied on comorbidity data alone or on small candidate-variant genetic scores, with inconsistent or incompletely reported performance. No study has combined a well-powered, multi-ancestry polygenic risk score (PRS) with environmental, reproductive, and symptom data in a single hybrid model. We developed and evaluated hybrid risk-prediction models integrating a genome-wide, multi-ancestry PRS with clinical and symptom data for endometriosis in the US-based All of Us Research Program. Methods: Among 69,376 participants (15,382 endometriosis cases, 53,994 controls) across six genetically inferred ancestry groups, we computed individual-level PRS values using PRS-CS weights derived from an independent, multi-ancestry GWAS. Five nested logistic regression, random forest, and XGBoost models progressively added age, ancestry, and within-ancestry genetic principal components (Model 1), environmental and reproductive factors (Model 2), symptom and comorbidity indicators (Model 3), all covariates combined (Model 4), and PRS x environment interactions (Model 5). Performance was assessed by AUROC in a held-out test set and 5-fold cross-validation, with class-weighted, Youden-optimized thresholds used for sensitivity, specificity, and predictive values; permutation importance identified top contributors. Pairwise AUROC differences were tested with a Holm-corrected DeLong-type test. Results: Discrimination improved from AUROC 0.63 (PRS, age, ancestry, principal components) to 0.72 for the full model, driven mainly by symptom and comorbidity data. XGBoost consistently outperformed logistic regression and random forest. The PRS ranked among the top individual predictors by permutation importance in nearly every model, alongside age, while genetic and demographic information alone gave only modest discrimination, and PRS x environment interactions did not improve on environmental factors alone. Threshold optimization yielded balanced sensitivity and specificity (~0.67/0.65) versus near-zero sensitivity at a default threshold. Conclusions: Combining the PRS with symptom and comorbidity data gave the best discrimination compared to solely a well-powered, multi-ancestry PRS as a predictor of endometriosis. This study clarifies both the promise and current limits of hybrid genetic-clinical prediction for endometriosis and points to symptom-based phenotyping, molecular subtyping, and external validation as priorities.
Thomas, J.; Pozdeyev, N.
Show abstract
Convolutional neural networks (CNNs) can classify thyroid nodules on ultrasound, yet published models are seldom available for independent testing, require machine learning expertise to develop and deploy, and are validated mostly on papillary thyroid carcinoma. Objective. To test whether an autonomous (agentic), no code artificial intelligence (AI) agent can develop a calibrated thyroid-nodule malignancy classifier, and to validate it internally and on an external cohort spanning multiple cancer histologies. Methods. This is a retrospective, computational diagnostic study with prespecified endpoints. A no code agent (Hugging Face ML Intern) autonomously reviewed data, selected and trained the model and calibrated probabilities, using the open source TN5000 dataset (3500 training, 500 validation, and 1000 test images). The trained ResNet 18 model was externally validated on 232 nodules from the University of Colorado, including follicular, medullary, oncocytic, and follicular variant of papillary carcinomas. Results. On the internal test set, an agentic AI model achieved AUROC 0.94 (95% CI, 0.920 - 0.953), sensitivity 0.90, and specificity 0.80. On external validation, agentic AI model achieved an AUROC of 0.90 (95% CI, 0.850 - 0.936), sensitivity of 0.92, and specificity of 0.68, negative predictive value of 0.96, and positive predictive value of 0.52, exceeding the performance of a previously published classifier on the same cohort (AUROC of 0.83). Conclusions. An agentic, no code AI workflow produced a calibrated, externally validated thyroid nodule classifier, supporting accessible, reproducible, and independently testable medical AI development. Prospective validation and local recalibration are required before clinical use.
Sanchez-Valle, J.; Zambrana, C.; Navarro-Martinez, A.; Costa, F. X.; Rocha, L. M.; Cirillo, D.; Violan, C.; Valencia, A.
Show abstract
Multimorbidity is the dominant clinical reality of primary care, yet the temporal dynamics governing when and how persistent comorbidity associations emerge remain poorly characterised. Most large-scale comorbidity studies adopt a single observation window after an index diagnosis, implicitly assuming that associations detectable at one year are equally detectable at five. Using 11 years of electronic health records from 5,821,197 individuals in Catalan primary care, we applied a matched cohort design across nine complementary follow-up windows, five cumulative (0-1 to 0-5 years) and four conditional (1-2 to 4-5 years), to 1,315 index diseases, identifying 144,030 significant directed comorbidity associations in the five-year network. We found that 60.1% of these associations required at least three years of follow-up and were undetectable in shorter-window analyses, demonstrating that observation window length is a primary determinant of which comorbidities can be observed. To organise this temporal heterogeneity, we introduce the biological clock of multimorbidity: a two-dimensional framework that positions ICD-10 disease categories according to their rates of cumulative signal attenuation and the persistence of conditional risk. This framework identifies four reproducible temporal patterns (episodic, chronic stable, chronic progressive, and transient-persistent) that are robust under bootstrap resampling, leave-one-disease-out sensitivity analysis, and alternative clustering approaches. The biological clock is systematically modulated by sex, with Blood/Immune and Musculoskeletal disorders showing the largest sex differences in temporal dynamics. Network analysis identified 19 disease "initiators" that generate broad downstream comorbidity burdens and 21 "sinks" representing convergent endpoints of multiple disease trajectories. Comparison with hospital-based Danish data from 6,909,676 individuals showed that shared associations were 2.7-fold enriched over chance expectation (hypergeometric test, p<10-300) and showed moderate concordance of effect sizes (Spearman {rho}=0.460), confirming that the comorbidity structure identified here reflects genuine, generalisable signal; nonetheless, only 3.6% of primary care associations were replicated in the hospital network, indicating that the two settings capture largely complementary segments of the disease co-occurrence landscape. Together, these findings establish the observation window length as a principal design parameter in EHR-based multimorbidity research and the biological clock as a framework for understanding how and over what timescale disease associations emerge, persist, and resolve.
KATUMBA, A. M.; Drakesmith, C. W.; Haynes, S.; Maynard, S.; Maharajan, V.; Erone, I.; Smith, M.; Shah, A.; Roy, N.; Bankhead, C.; Stanworth, S. J.
Show abstract
Background Iron deficiency (ID) is a readily treatable condition once identified. Ferritin is the primary diagnostic marker, but cut-offs vary and inflammation complicates interpretation in patients with long-term conditions (LTCs). Aim To describe ferritin distribution and the prevalence of threshold-defined low ferritin in adults with and without LTCs in primary care. Design and setting Cross-sectional observational study using routinely collected electronic health records from a national primary care database in England (1st January 2015 to 31st December 2021). Method Adults with >1 ferritin test in Clinical Practice Research Datalink (CPRD) Aurum were included. LTCs were identified using validated primary-care code lists. Outcomes included ferritin distribution and threshold-defined ID prevalence using World Health Organization (WHO) (<15 ug/L; <70 ug/L if inflammation) and National Institute for Health and Care Excellence (NICE) (<30 ug/L) cut-offs, stratified by sex and, in women, by age <50 versus >=50 as a proxy for menopausal status. Results 4,489,594 individuals were included; 55% (n=2,469,882) had >1 LTC. Ferritin was lowest in women <50 and in LTCs characterised by impaired absorption or blood loss (coeliac disease, inflammatory bowel disease). Among women <50 with an LTC, 80% had ferritin <70 ug/L versus 47% <30 ug/L, leaving 33% in the 30 to 70 ug/L range potentially missed by standard cut-offs; equivalent figures were 28% in women >=50 and 17% in men. Conclusion Threshold-defined low ferritin is very common across LTCs and disproportionately affects women, particularly those under 50. Condition-specific, inflammation-adjusted ferritin thresholds may improve detection, management, and equity in primary care.
Ayati, A.; Onal, G.; Sur, A.; Azzam, S.; Wang, B.; Rudrapatna, V. A.
Show abstract
Objective: Erythropoietic protoporphyria (EPP) is a rare photodermatosis marked by multi-year diagnostic delays. We developed and externally validated machine learning models to identify patients with EPP earlier from longitudinal electronic health record (EHR) data and estimate undiagnosed disease burden. Materials and Methods: In a retrospective case-control study at two San Francisco health systems, an academic referral center (UCSF) and a safety-net hospital (ZSFG) we identified 74 confirmed EPP cases using combined diagnostic coding, biochemical criteria, and specialty chart review. Symptom-enriched controls were sampled at a 40:1 ratio. Longitudinal diagnoses, laboratory results, medications, procedures, and encounters preceding the outcome date were modeled with a gradient-boosting classifier (CatBoost) and a state-space sequence model (MAMBA). The best model was deployed across the UCSF population and externally validated at ZSFG without retraining. Results: On the UCSF held-out test set (n=1,865; 43 cases), MAMBA outperformed CatBoost (AUC ROC 0.91 vs 0.89; average precision 0.42 vs 0.27; precision 65% vs 20%), flagging cases a median of 229 days before documented diagnosis. Deployed across 297,967 symptom-compatible patients, it identified 310 high-risk individuals, implying a prevalence approaching genetic estimates. External validation at ZSFG showed attenuated performance (AUC ROC 0.72; average precision 0.10) while preserving early detection (median 264 days). Discussion: A sequence model integrating temporal EHR signals detected EPP months before clinical recognition, corroborating genetic evidence of substantial underdiagnosis. Cross-site attenuation reflects population and documentation differences and underscores the need for local recalibration. Conclusion: Longitudinal EHR-based machine learning can shorten EPP diagnostic delay and prioritize patients for confirmatory testing, supporting proactive rare-disease case finding.
Kostan, H.; Krishnan, V.
Show abstract
Persons with epilepsy (PWE) experience high rates of psychiatric comorbidity, yet the population-scale pharmacoepidemiology of psychiatric medication (PM) use alongside antiseizure medications (ASMs) has not been previously characterized. Using Epic Cosmos, a federated electronic health record network spanning >300 million patients across >2,000 hospital systems, we examined patterns of ASM and PM co-prescriptions in PWE (ICD- 10 G40.x) every year between 2018-2025. PM prescriptions were similarly assessed in patients with asthma (J45.x). We found that despite the introduction of several newer ASMs, the overall prescribing landscape remained stable, with little change in the relative use of individual ASMs over time. Compared with asthma patients, PWE were more likely to receive prescriptions for opiates, antidepressant and antipsychotic medications across the age spectrum. 17-year-old or younger PWE were more likely to receive ADHD/stimulant medications, whereas adults and older adults exhibited a shift toward cognitive enhancing agents. We did not observe a preponderance of PM co- prescribing with any specific ASM or ASM class. Together, these results provide a population-scale, age-specific survey of psychiatric medication co-prescriptions in epilepsy, establishing a framework to monitor surrogate markers of psychiatric comorbidity and to support pharmacovigilance of potential drug-drug interactions.
Olshvang, D.; Harris, C. W.; Chellappa, R.; Parikh, C.; Santhanam, P.
Show abstract
Background Long-horizon kidney trajectory prediction in type 2 diabetes mellitus (T2DM) is usually reported as a point estimate or event risk, although clinical decision-making also depends on whether an individual prediction is reliable. We developed an uncertainty-aware model for 48-month estimated glomerular filtration rate (eGFR) decline and tested whether conformal interval width provides a clinically structured, patient-level signal of prediction reliability. Methods We performed a secondary prognostic modeling analysis of Action to Control Cardiovascular Risk in Diabetes (ACCORD) participants with baseline and 48-month eGFR (n=6,853). The outcome was annualized eGFR change, calculated as 48-month minus baseline eGFR divided by four years. The primary baseline feature set excluded serum creatinine since eGFR is creatinine-derived, and also excluded urine biomarkers. Random forest, gradient boosting, penalized linear models, and XGBoost were compared using fixed training, calibration, and test partitions. Split and locally adaptive conformal intervals were evaluated by empirical coverage and interval width. Interval-width analyses were repeated after conditioning on baseline eGFR. Results The best primary model was random forest (R2=0.382, MAE=3.271 mL/min/1.73m2). Split 90% conformal intervals achieved empirical coverage of 0.917. Locally adaptive 90% intervals achieved empirical coverage of 0.909 with mean width 13.759 mL/min/1.73m2. In unadjusted analyses, wider intervals were associated with larger errors and more rapid decline. After interval-width quintiles were assigned within baseline-eGFR strata, wider intervals remained associated with realized prediction error (annual adjusted increase, 0.151 mL/min/1.73m2 per quintile). Beyond baseline eGFR, wider intervals were associated with younger age, female sex, higher HbA1c, higher triglycerides, and higher systolic blood pressure. Conclusions Baseline clinical variables predicted 48-month eGFR decline with good long-horizon performance in ACCORD, even after excluding serum creatinine and urine biomarkers from the primary model. Conformal prediction provided calibrated patient-specific intervals, and interval width behaved as an informative reliability phenotype rather than a random modeling artifact. These findings support a novel uncertainty-aware framing of kidney trajectory prediction in which rapid and uncertain decline can be identified from baseline clinical data.
Chen, M.; Huang, Y.; Yu, R.; Xie, Y.; Chen, F.; Huang, J.; Zhao, J.; Ma, Z.; Ma, Z.; Jiang, L.
Show abstract
Background: Hearing loss is a potentially modifiable risk factor for brain health, but whether it acts as a causal lever remains unclear. Methods: We constructed an ear-disease comorbidity network from NHANES 2011-2020 (N=18,939, 16 nodes, 62 edges), performed bidirectional Mendelian randomization (MR) across 24 exposure-outcome pairs, and triangulated evidence with longitudinal data from CHARLS (N=17,101). Results: Subjective hearing symptoms (prevalence 6.1%) occupied hub positions in the comorbidity network, whereas objective hearing impairment (8.0%) was sparsely connected. All forward MR estimates were null after multiple-testing correction (IVW P>.05 for 9 of 9 pairs). Reverse MR showed one nominally significant association (cognition to objective hearing beta=-0.15, P=.013) that did not survive correction. Longitudinal analysis yielded HR=1.57 (P=.00004) for subjective hearing symptoms predicting incident depression. Conclusions: Perceived hearing symptoms organize the ear-disease comorbidity network but are not a causal lever for brain health. These findings support a "flag, not lever" framework: subjective hearing symptoms warrant clinical attention as markers of systemic multimorbidity rather than intervention targets for dementia prevention.
Oliveira, J. F.; Alencar, A. L.; Coutinho, E. R.; Borges, D. G. F.; Filho, F. M. H. S.; Santos-Silva, R.; Tavares Veras Florentino, P.; Cunha, M. C. S. L.; Marcilio, I.; Pereira Ramos, P. I.; Andrade, R. F. S.; Barral-Netto, M.
Show abstract
Background: Evaluating outbreak detection models is a key component of syndromic surveillance. However, balancing timeliness, predictive performance, and local surveillance constraints remains a major challenge. We developed and assessed whether stacking ensemble approaches, which integrate multiple outbreak detection methods, can improve the timeliness and predictive performance of influenza-like illness (ILI) surge detection. Methods: We developed a two-stage stacking ensemble framework to detect early warning of ILI surges in city-level Primary Health Care encounter time series from Brazil (2022 to 2025). Epidemic thresholds were defined using the Moving Epidemic Method (MEM). In the first stage, multiple outbreak detection models (ODMs) generated warnings of unusual ILI activity. In the second, these warnings were then used as inputs to three supervised meta-classifiers: Logistic Regression, Extreme Gradient Boosting (XGB), and a Multi-layer Perceptron (MLP). For comparison, a Majority Voting (MV) aggregation is also examined. Timeliness, sensitivity, specificity, positive and negative predictive values are evaluated to measure each model's ability to anticipate epidemic periods of varying intensity in 2025. Robustness was further assessed using simulated outbreak scenarios with varying magnitudes and durations. Findings: We identified 5,765 ILI surge onsets across 5,365 Brazilian municipalities in 2025. Compared with individual ODMs and MV, stacking ensemble meta-classifiers anticipated up to 33% of surge onsets three weeks in advance (an average improvement of 15 percentage points) while reducing missed detections to <10%. They achieved sensitivity >90%, while maintaining balanced specificity >80%, PPV >65%, and NPV >99%. Improvements were greatest for very high-intensity surges, with missed detections reduced by more than half compared with individual ODMs. In simulated outbreak scenarios, the MLP and XGB classifiers remained robust despite being trained on fewer than half of all simulated surge events, consistently outperforming individual detection methods and simpler integration approaches. Interpretation: We provide a practical framework for integrating complementary ODMs into a single, robust early warning decision. By improving both timeliness and predictive performance without requiring additional surveillance data or resources, this approach offers a scalable methodological upgrade for syndromic surveillance systems and supports more reliable public health decision-making. Funding: The Rockefeller Foundation (award 2023 PPI 007 to MB-N); Brazilian National Research Council - CNPq (408775/2024-6); MB-N, PIPR, RFSA are CNPq fellows.
Holmes, J. P.; Zutautas, K. B.; Sisnett, D. J.; Hayati, D.; Bougie, O.; Lessey, B. A.; Tayade, C.
Show abstract
Endometriosis (EM) is a heterogeneous, gynecological inflammatory disease affecting over 200 million individuals worldwide, yet the mechanisms underlying lesion establishment, progression, and recurrence remain incompletely understood. Small extracellular vesicles (sEVs) mediate intercellular communication through the transfer of proteins, lipids, and nucleic acids reflective of their cellular origin; however, stage- and tissue-specific sEV signatures remain poorly defined. Here, we characterized the molecular and functional landscape of EM-derived sEVs across disease stages and biological sources. sEVs isolated from eutopic endometrium, ectopic lesions, peritoneal fluid, and plasma from mild- and severe-stage EM patients and healthy controls were analyzed by surface marker profiling, proteomics, lipidomics, and integrated multi-omics, with functional effects assessed in human uterine microvascular endothelial cells. sEV composition varied by disease stage and sample type, with EM lesion-derived sEVs demonstrating stage-dependent loss of epithelial-associated markers and enrichment of immune-associated signatures, while EM plasma-derived sEVs exhibited altered adhesion- and platelet-associated profiles. Integrated multi-omics identified coordinated programs associated with immune adaptation, extracellular matrix organization, epithelial remodeling, vascular signaling, oxidative stress, and metabolic adaptation. Functionally, sEVs derived from severe endometriotic lesions exhibited enhanced uptake and mitochondrial localization in endothelial cells and promoted angiogenic activity. Our findings establish sEVs as dynamic mediators of EM disease progression and demonstrate that integrated sEV profiling provides a framework for understanding EM heterogeneity and identifying candidate biomarkers and therapeutic targets.
Hanzlikova, Z.; Styk, J.; Pös, O.; Biro, O.; Bokorova, S.; Lukyova, L.; Sitarcik, J.; Sladecek, T.; Krampl, W.; Meszaros, A.; Hunyadi, P.; Mate, S.; Egeto, A.; Rigo, J.; Sedlackova, T.; Radvanszky, J.; Budis, J.; Szemes, T.
Show abstract
Background: Despite advances in circulating tumor DNA analysis, reliable detection of oncological disease from ultra-low coverage whole genome sequencing (ulcWGS) remains challenging, particularly at low tumor fractions. This study leverages cell-free DNA (cfDNA) characteristics to develop and evaluate a robust, integrative binary predictive model for ovarian cancer (OC) status screening. OC represents a growing global burden and is often diagnosed at advanced stages due to the lack of specific early symptoms and effective screening strategies, highlighting the need for sensitive and broadly applicable early detection approaches.Methods and Findings: We analyzed plasma cfDNA from OC patients (N = 85) and cancer-free controls (N = 41) using ulcWGS (~1x). Participation in the study was voluntary, and all participants provided written informed consent before any study-related procedures under study approval No. 16119-8/2022/EUIG. Within an integrated workflow combining standardized laboratory processing, bioinformatic pipelines, and machine learning (ML), we extracted 21 features capturing copy number variations (CNVs) and fragmentomic characteristics to identify complementary signatures distinguishing OC from controls. Predictive models were developed using XGBoost with hyperparameter optimization and evaluated on an independent test set (n = 25% of the cohort). A dual-threshold classification strategy was applied to define an uncertainty zone and optimize screening performance.CNV-derived and fragmentomic features assessed in exploratory analysis on the training-validation set showed moderate discriminative power (AUC 0.569 - 0.946) but substantial overlap between groups. On the test set, the CNV-only model achieved an AUC of 0.855 (sensitivity 85%, specificity 50%), while the fragmentomics-only model reached an AUC of 0.8825 (sensitivity 95%, specificity 30%). Both feature domains captured complementary aspects of tumor-derived cfDNA, with fragmentomics favoring sensitivity and CNV-derived metrics improving specificity. Integration of both feature classes improved performance, yielding an AUC of 0.900, sensitivity of 85.00%, and specificity of 90.00%. SHAP analysis confirmed contributions from both feature types without a single dominant predictor.Conclusions: We present an integrative cfDNA framework for OC detection based on ulcWGS that combines CNV and fragmentomic signals to improve diagnostic performance over single-feature approaches. By enabling robust detection of tumor-associated patterns at ultra-low sequencing depth, this approach demonstrates that meaningful cancer discrimination can be achieved without reliance on deep sequencing. This highlights the potential of cost-effective and scalable liquid biopsy strategies for population-level cancer screening their integration into personalized and preventive oncology. Keywords: Liquid biopsy, ovarian cancer, ultra-low coverage whole genome sequencing, cell-free DNA, cell-free tumor DNA, cancer detection, copy number variations, insert size, fragmentomics, machine learning
Katarynczuk, K.; Stachowiak, A.; Piorkowska, N. J.; Ostromecki, A.; Franik, G.; Bizon, A.
Show abstract
Background: Machine-learning models for polycystic ovary syndrome (PCOS) and other conditions frequently report near-perfect diagnostic performance, but retrospective datasets assembled from routine clinical practice can encode diagnostic-group membership in how data were acquired rather than in disease biology, and this acquisition-related information can be indistinguishable from genuine clinical signal under conventional validation. Objective: To determine, using a real-world PCOS cohort as a case study, whether high classification performance reflected clinically meaningful information or artifacts of data provenance, schema structure, and measurement-acquisition workflow, and to develop a generalizable audit framework for detecting such artifacts in retrospective medical machine learning. Methods: We analyzed 1,331 retrospective records (1,286 PCOS, 45 controls) from a single endocrine-gynecology database. A layered acquisition-bias framework compared classification performance using (i) raw and harmonized missingness patterns alone, (ii) measured values with and without explicit missingness indicators, and (iii) ascertainment-balanced feature sets with and without age. Logistic regression and random forest were evaluated using repeated stratified cross-validation, bootstrap resampling, label-permutation testing, and calibration analysis, and the framework was validated against a semi-synthetic experiment with known ground truth. Results: Diagnostic status was perfectly predicted (ROC-AUC = 1.000) from missingness patterns alone, before any clinical value was examined, and this persisted after semantic harmonization of duplicated source columns. Performance declined progressively as acquisition-sensitive information was removed, from near-ceiling in raw and harmonized value models to a mean ROC-AUC of approximately 0.80-0.82 in the most restrictive ascertainment-balanced, age-excluded representation. The semi-synthetic experiment reproduced this pattern under known data-generating conditions, confirming that harmonization removes schema-fragmentation artifacts but not workflow-driven acquisition bias. Conclusions: Apparent diagnostic performance in this cohort was substantially attributable to diagnostic workflow and data-acquisition structure rather than to a stable, transportable biological signal. The layered audit framework generalizes beyond PCOS and offers a practical tool for detecting acquisition-related leakage in retrospective clinical machine-learning studies.
Dao, V. N.; Nguyen, P. T.; Tran, T. N.; Nguyen, N. H.; Tang, H.-S.; Boni, M. F.; Giang, H.; Phan, D. M.
Show abstract
Non-invasive prenatal testing (NIPT) was initially developed to detect chromosomal abnormalities in fetuses through the analysis of cell-free fetal DNA in maternal blood. Recent advancements have expanded NIPT's applications to include the detection of viral infections during pregnancy. However, interpreting pathogen-derived cell-free DNA (cf-DNA) remains clinically complex. This study explores the clinical relevance of hepatitis B virus (HBV) cf-DNA using a dataset of approximately 500,000 NIPT visits and an independent validation cohort of 582 pregnant women (40 HBV-infected), aligned with HBV epidemiology from both population and individual perspectives. Our analysis reveals that HBV cf-DNA is a strong biomarker of high viral infectivity rather than a general marker of infection, suggesting its potential to identify pregnant women at heightened risk of vertical transmission by the end of the first trimester. Additionally, HBV-positive women showed a small but consistent reduction in fetal fraction relative to HBV-negative women across gestational weeks 9 - 17, an association compatible with an early effect of HBV on the placental contribution to cell-free DNA, although the observational design and unmeasured maternal covariates preclude causal inference.
Kathuria, Y.; Miller, K.; Selden, E. B.; Gallagher, W. J.; Capan, M.
Show abstract
Patients diagnosed with type 2 diabetes (T2D) are at increased risk of developing cardiovascular disease (CVD), the leading cause of morbidity and mortality in this population. Early detection and glycemic control within the first year after diagnosis reduce CVD risk. However, gaps remain in how to operationalize early detection of T2D using Electronic Health Record (EHR) data and quantify its relationship with subsequent CVD risk using longitudinal observations. We developed a probabilistic graph model to analyze the interdependencies between early detection of T2D, post-diagnosis glycemic control, and CVD occurrence. Using a temporally structured Bayesian Network (BN) learned from EHR data of 9,450 primary care patients between 2017 and 2023, we quantified probabilistic dependencies between demographics, diagnostic delay surrogates, glycemic control, and post-diagnosis CVD occurrence. Percentile based thresholds defined risk groups, where individuals with predicted probabilities in the bottom decile ([≤] 10th percentile) were classified as low risk, and those in the top decile ([≥] 90th percentile) as high risk. Results demonstrated heterogeneity in predicted risks across glycemic and cardiovascular outcomes. Predicted probability of developing CVD within the first year after T2D diagnosis ranged from a mean of 5.2% in the low-risk group to 28.9% in the high-risk group, while predicted probabilities of mean Hemoglobin A1c (HbA1c) [≥] 8% during the first year post-diagnosis ranged from 1.6% in low-risk to 55.1% in high-risk group. Patients with HbA1c at diagnosis [≥] 8% had higher predicted probabilities of first-year post-diagnosis mean HbA1c [≥] 8% (53.3% vs. 1.9%) and high HbA1c coefficient of variation (18.7% vs. 3.1%) compared with those with HbA1c [≤] 6.5%. Incorporating early clinical outcomes refined later risk predictions, with long-term CVD risk reaching 33.5% among high-risk individuals. The proposed model achieved predictive performance comparable to conventional machine learning approaches while providing interpretable relationships for risk stratification in primary care populations.
Pardal-Refoyo, J. L.; Zapatero-Sanchez, I.
Show abstract
Background Bethesda III thyroid nodules remain diagnostically indeterminate, and intraoperative frozen section is used selectively to support surgical decision-making. Its value in this specific cytological category is uncertain because a malignant result may be highly specific while non-malignant and non-definitive results may fail to exclude cancer. Objective To estimate the sensitivity and specificity of intraoperative frozen section for detecting malignancy in thyroid nodules with preoperative Bethesda III cytology. Methods The protocol was prospectively registered in PROSPERO (CRD420261416683). A systematic review was conducted in PubMed, Embase, Web of Science, Europe PMC, and the Cochrane Library. Studies were eligible when they reported a separable Bethesda III cohort, intraoperative frozen-section findings, and final histopathology. Frozen section was classified as positive only when malignancy was reported. Benign, suspicious, indeterminate, deferred, inconclusive, and follicular-pattern results were classified as non-malignant. Study-level 2 x 2 tables were synthesised with random-effects logit models. Because all studies reported zero false-positive results, a full bivariate model with freely estimated covariance was not identifiable; a pseudo-bivariate HSROC approximation was therefore used. QUADAS-2 was used for risk-of-bias assessment. Results Ten studies comprising 1,069 Bethesda III patients or nodules were included. The pooled sensitivity was 43.1% (95% CI, 21.3-67.9) and the pooled specificity was 98.8% (95% CI, 97.0-99.5). Sensitivity was highly heterogeneous (Q = 83.45, I2 >90%, T2= 2.15), whereas specificity showed negligible between-study variance. Excluding Mao 2023 reduced pooled sensitivity to 35.6% (95% CI, 23.8-49.5) and reduced sensitivity heterogeneity to moderate levels (Q = 12.44, I2 = 36%, T2 = 0.25); specificity remained 98.7% (95% CI, 96.7-99.5). The approximate HSROC AUC was 0.89 including Mao 2023 and 0.87 excluding Mao 2023, but these values were driven largely by the uniformly high specificity. Conclusions Intraoperative frozen section in Bethesda III nodules has excellent specificity but limited and variable sensitivity. It is more suitable as a rule-in test than as a rule-out test. A non-malignant or non-definitive result should not be used alone to exclude malignancy or determine the extent of thyroid surgery.